Papers with student networks
TelME: Teacher-leading Multimodal Fusion Network for Emotion Recognition in Conversation (2024.naacl-long)
Copied to clipboard
| Challenge: | Emotion Recognition in Conversation (ERC) aims to identify emotions expressed by participants at each turn within a conversation. |
| Approach: | They propose a Teacher-leading Multimodal fusion network for ERC that integrates cross-modal knowledge distillation to transfer information from a lan- guage model acting as the teacher to non- verbal students. |
| Outcome: | The proposed model achieves state-of-the-art in a multi-speaker conversation dataset for ERC. |
Annealing Knowledge Distillation (2021.eacl-main)
Copied to clipboard
| Challenge: | Knowledge distillation (KD) is a powerful model compression technique for deep neural networks. |
| Approach: | They propose a method to feed the rich information provided by teacher’s soft-targets incrementally and more efficiently by annealing the teacher output incrementally. |
| Outcome: | The proposed method can be used on image classification and NLP language inference tasks with BERT-based models on the GLUE benchmark. |
Improving the Robustness of Distantly-Supervised Named Entity Recognition via Uncertainty-Aware Teacher Learning and Student-Student Collaborative Learning (2024.findings-acl)
Copied to clipboard
| Challenge: | Named Entity Recognition (NER) methods require a substantial quantity of high-quality annotation for training models. |
| Approach: | They propose a method to reduce the number of incorrect pseudo labels in self-training . they propose 'uncertainty-aware teacher learning' and 'student-student collaboration' |
| Outcome: | The proposed method is superior to state-of-the-art DS-NER denoising methods. |
Continuation KD: Improved Knowledge Distillation through the Lens of Continuation Optimization (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods for knowledge distillation (KD) do not mitigate the noise in the teacher’s output: modeling the noisy behaviour of the teacher can distract the student from learning more useful features. |
| Approach: | They propose a method that optimizes the highly non-convex KD objective by starting with the smoothed version of this objective and making it more complex as the training proceeds. |
| Outcome: | The proposed method achieves state-of-the-art performance on NLU and computer vision tasks. |